The problem: how do you train 10 million parameters?
In Lesson 4.1 you built a network with three layers. That small network already had thousands of individual weight parameters. Real-world networks today have millions or billions. You cannot adjust each weight by hand or by trial and error. You need an algorithm that can figure out exactly how to nudge every single parameter to make the network better.
That algorithm is backpropagation, short for "backward propagation of errors." It was popularised in a landmark 1986 paper by David Rumelhart, Geoffrey Hinton, and Ronald Williams published in the journal Nature. While the mathematical ideas had been developed independently by others in the preceding years, the 1986 paper made the algorithm accessible and demonstrated that it could teach hidden layers to form useful internal representations, which settled a long-standing debate about whether multi-layer networks could be trained at all.
A chef makes a dish and it tastes wrong. To fix it, the chef needs to know which ingredient was the problem and by how much. Backpropagation does exactly this for neural networks. After a bad prediction, it works backwards through every layer and calculates how much each weight contributed to the error. Then it adjusts each weight in the direction that reduces that contribution. It assigns blame precisely and efficiently.
How backpropagation works
Backpropagation has two phases. First, the forward pass: data flows forward through the network and a prediction is made. Second, the backward pass: the error is computed, and then it is propagated backwards through the network layer by layer, updating weights along the way.
The forward pass produces a prediction (gold arrows). The loss function compares that prediction to the true label and outputs a single error number. The backward pass (dashed red) flows that error gradient back through every layer, computing how much each weight should change. Every weight gets updated by a tiny amount in the direction that reduces the error.
The mathematics of backpropagation relies on the chain rule from calculus. For each weight, you compute: "if I increase this weight by a tiny amount, how much does the overall loss change?" This rate of change is the gradient. You then move the weight slightly in the direction that decreases the loss. You do not need to know the calculus to use neural networks, but understanding that gradients are the mechanism of learning is important for diagnosing problems.
The vanishing gradient problem
When networks become very deep, a serious problem emerges: gradients can become vanishingly small as they travel backwards through many layers. Each layer multiplies the gradient by the derivative of its activation function. For sigmoid and tanh activations, this derivative is at most 0.25. Multiply a small number by 0.25 twenty times and it becomes essentially zero.
When gradients vanish, the weights in the early layers of the network receive almost no signal. They barely update. The network struggles to learn anything in its first layers. This was the key reason that very deep networks were considered impractical for much of the 1990s and 2000s.
Three developments largely solved the vanishing gradient problem:
1. ReLU activation: For positive inputs, the derivative of ReLU is exactly 1. No shrinkage. Signals can propagate back through many layers without fading, making deep networks trainable.
2. Residual connections (ResNets, 2015): Kaiming He and colleagues at Microsoft Research introduced "skip connections" that let the gradient bypass entire blocks of layers. A residual connection adds the input of a block directly to its output: output = F(x) + x. This gives the gradient a shortcut path that avoids the problem entirely. ResNets enabled networks 100 or more layers deep to train successfully.
3. Better weight initialisation: Initialising weights with the wrong scale can cause gradients to explode or vanish even before training begins. Methods like He initialisation (for ReLU networks) and Glorot/Xavier initialisation (for tanh networks) set the starting scale of weights so that signals are preserved as they pass through the network.
Loss functions: measuring what the network gets wrong
Before backpropagation can compute any gradients, it needs a number to differentiate: the loss. The loss function converts the difference between prediction and truth into a single number that can be minimised. Choosing the right loss function for your problem is not optional.
| Problem type | Loss function | Why |
|---|---|---|
| Binary classification | Binary cross-entropy | Measures the gap between a predicted probability and a 0/1 label. The natural choice when your output is a sigmoid probability. |
| Multi-class classification | Categorical cross-entropy | Generalises binary cross-entropy to multiple classes. Used with a softmax output layer. Each class gets a probability and the loss rewards high probability on the correct class. |
| Regression | Mean Squared Error (MSE) | Penalises large errors more heavily than small ones. The standard choice for continuous outputs. Use Mean Absolute Error if large outliers should be treated more gently. |
Optimisers: how to descend the gradient
Knowing the gradient tells you which direction to move each weight. The optimiser decides how far to move and how. Vanilla gradient descent (updating all weights once per pass through the entire dataset) is rarely used in deep learning because it is too slow and can get stuck.
| Optimiser | Key idea | When to use |
|---|---|---|
| SGD | Update weights after each mini-batch. Faster than full-batch gradient descent but noisier. Adding momentum smooths the updates and helps escape local minima. | When you want fine-grained control and are willing to tune the learning rate manually. Often best for CNNs with a careful schedule. |
| Adam | Adapts the learning rate for each parameter individually, using estimates of the first and second moments of the gradient. Introduced by Kingma and Ba in 2014. | The default starting choice for most networks. Works well with little tuning. Slightly higher memory use than SGD. |
| AdamW | Adam with decoupled weight decay regularisation. Fixes a subtle issue in how Adam applies L2 regularisation. | The default for training large language models and transformers. Generally preferred over Adam when regularisation matters. |
The learning rate controls how large each weight update step is. Too large and the updates overshoot the minimum and the loss oscillates or diverges. Too small and training takes an impractically long time. A common practical approach is to start with a moderate learning rate and then reduce it (using a learning rate schedule) as training progresses. The Keras default learning rate for Adam is 0.001, which is a reasonable starting point for most problems.
Reading the training curve
When you train a neural network, Keras tracks the loss on both the training set and the validation set at each epoch. Plotting these two curves over time is the single most useful diagnostic tool in deep learning. The shape of the curves tells you exactly what is happening.
Plot both curves after every training run. If they diverge (training keeps improving while validation worsens), you are overfitting. Stop training earlier or add regularisation. If both stay high, your model lacks capacity or needs more training time.
Regularisation: preventing overfitting in neural networks
Deep networks have enormous capacity. Left unconstrained on a small dataset, they will memorise training examples rather than learning generalisable patterns. Two regularisation techniques are standard practice in neural networks.
from tensorflow.keras import layers, callbacks import tensorflow as tf model = tf.keras.Sequential([ layers.Dense(128, activation='relu'), layers.BatchNormalization(), # normalise layer outputs layers.Dropout(0.3), # randomly zero 30% of neurons layers.Dense(64, activation='relu'), layers.BatchNormalization(), layers.Dropout(0.3), layers.Dense(1, activation='sigmoid') ]) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) # Early stopping: stop if val_loss doesn't improve for 10 epochs early_stop = callbacks.EarlyStopping( monitor='val_loss', patience=10, restore_best_weights=True # revert to the best epoch when stopping ) history = model.fit( X_train, y_train, epochs=200, batch_size=32, validation_split=0.15, callbacks=[early_stop], verbose=0 ) print(f"Stopped at epoch {len(history.history['loss'])}")
With restore_best_weights=True in the EarlyStopping callback, Keras automatically reverts the model to its best state before the validation loss started worsening. You ask it to train for 200 epochs but it might stop at epoch 47, having found the best generalisation well before the end.
"The practical utility of the various tricks used in training neural nets is often underestimated. Choosing good activations, initialisations, and optimisers matters as much as choosing the architecture."
Andrej Karpathy, Stanford CS231n lecture notes (widely read AI teaching resource)